Back

NAR Genomics and Bioinformatics

Oxford University Press (OUP)

Preprints posted in the last 90 days, ranked by how well they match NAR Genomics and Bioinformatics's content profile, based on 242 papers previously published here. The average preprint has a 0.17% match score for this journal, so anything above that is already an above-average fit.

1
Development of the Mitochondrial Base Editor Analysis Package (MitoBEAP).

Mutti, C. D.; Nash, P.; Silva-Pinheiro, P.; Minczuk, M.; Van Haute, L.

2026-06-05 bioinformatics 10.64898/2026.06.02.729539 medRxiv
Top 0.1%
17.8%
Show abstract

For many years, the genetic manipulation of mitochondrial DNA was largely hampered by inefficient delivery of nucleic acids to mitochondria. However, the development of mitoCBEs, such as mitochondrial cytosine base editors (DdCBEs), which catalyse C*G-to-T*A conversions, and more recently, mitoABEs, such as transcription-activator-like effector (TALE)-linked deaminases (TALEDs) enabling A*T-to-G*C conversion, has transformed this field. Generally, mitochondrial base editors exhibit high on-target efficiency and are straightforward to design and use. Nonetheless, unintended off-target effects cannot be overlooked and should be assessed consistently with each experiment, which can be challenging without specialised bioinformatic expertise. Here, we introduce Mitochondrial Base Editor Analysis Package (MitoBEAP), which, to our knowledge, is the first R package specifically designed to analyse next-generation sequencing data from base-edited mtDNA samples. The package facilitates the analysis of potential off-target effects, offers multiple visualisation options, and allows customisation of graphics and thresholds for calculations. As a proof of concept, this study demonstrates how MitoBEAP can be utilised to measure the efficiency of DdCBE treatment targeting human 12S rRNA, as well as to identify potentially harmful off-target conversions across the mtDNA.

2
Gene-specific exponent-corrected normalization for library size in bulk RNA-seq

Yin, R.; Li, D.; Zong, W.; Ketchesin, K. D.; Seney, M. L.; McClung, C. A.; Baldoni, P. L.; Tseng, G. C.

2026-07-09 bioinformatics 10.64898/2026.07.04.736167 medRxiv
Top 0.1%
14.9%
Show abstract

Correcting for library size is an essential step in bulk RNA-seq analyses, as differences in sequencing depth across samples can obscure biological signal with technical noise. While numerous normalization methods and model-based strategies have been proposed, we demonstrate here that library size-normalized counts and differential expression results obtained from such widely adopted approaches often remain strongly correlated with library size in large-scale RNA-seq experiments. Through a systematic analysis of over 100 publicly available GEO and TCGA RNA-seq datasets with raw count data, we show that library size association is observed for a substantial proportion of genes even after state-of-the-art library size correction approaches recommended by leading normalization tools. To address this issue, we propose gecco, a gene-specific exponent-corrected normalization method for RNA-seq counts that incorporates library size directly into the statistical framework via a gene-specific correction term, rather than applying a uniform adjustment factor across all genes. This formulation generalizes existing normalization approaches and yields normalized counts that are free of residual library size effects. Using both simulation studies and real large-scale RNA-seq datasets, we show that our method mitigates library size bias while preserving biological signal across a range of parameter settings. We further demonstrate that our approach leads to higher detection accuracy and more biologically meaningful pathway enrichment results in downstream differential expression and rhythmicity analyses without compromising false discovery rate control. Our method is implemented in R and is fully compatible with the widely used differential expression analysis methods DESeq2 and edgeR.

3
KozakExplorer: an interactive framework for genome-wide Kozak sequence analysis

Cokelaer, T.; Santi, A. M. M.; Pipoli da Fonseca, J.; Spaeth, G. F.

2026-06-26 bioinformatics 10.64898/2026.06.25.734688 medRxiv
Top 0.1%
14.9%
Show abstract

Translation initiation signals shape gene expression across all domains of life. In eukaryotes, nucleotide constraints surrounding the start codon are commonly described by the Kozak Consensus Sequence (KCS), whereas in bacteria and archaea, initiation frequently involves Shine--Dalgarno ribosome-binding motifs. Although these signals have been extensively characterized in model organisms, their large-scale diversity and evolutionary distribution remain incompletely explored. We present KozakExplorer, a reproducible framework for quantitative and comparative analysis of translation initiation contexts from genome assemblies and annotations. The software performs strand-aware extraction of start codon environments from FASTA and GFF3 files and applies information-theoretic metrics---including Kullback--Leibler (KL) divergence and information content (IC)---to measure positional nucleotide constraints relative to a background model. Derived summary statistics (Kozak Strength Index [KSI], maximum information content, peak position) convert motif patterns into interpretable per-genome signatures suitable for cross-species comparison. Our primary analysis covers 2,282 eukaryotic reference genomes, producing a standardized dataset of translation initiation metrics. Dimensionality reduction via t-SNE on per-position KL divergence, information content, and motif nucleotide frequencies reveals a structured eukaryotic KCS landscape with kingdom-level clustering and continuous variation in signal strength. A dedicated case study of 216 Apicomplexa genomes shows genus-level structure consistent with host range and phylogeny. An extended analysis across 25,344 reference genomes (22,253 bacteria, 809 archaea) places eukaryotic patterns in a global comparative framework, revealing transitions between sharply localized Kozak motifs and distributed Shine--Dalgarno-type signatures. Implemented within the open-source Sequana ecosystem, KozakExplorer is distributed as a Python module and an interactive web application that accepts local annotated assemblies, GenBank records, or NCBI RefSeq accessions, and exports all computed metrics, embeddings, and coordinates for downstream comparative and evolutionary genomics.

4
Trinucleotide Distribution, Symmetry Elements and Formulation of Mirror Symmetry Index for G4 Motifs

Arya, A.; Datta, B.

2026-07-05 bioinformatics 10.64898/2026.07.05.736592 medRxiv
Top 0.1%
12.8%
Show abstract

Symmetry elements in nucleic acids are most strongly correlated with sites of biological function; however, their relevance to non-canonical structures remains underexplored. In this study, we demonstrate the presence and significance of trinucleotide symmetry elements within G-quadruplex (G4) motifs. Our central hypothesis is that the intra-strand mirror symmetry of trinucleotides has been evolutionarily selected to facilitate G4 formation builds on the established sequence-structure association of G-quadruplexes and the natural symmetry law governing nucleotide insertion during genome evolution. Using a conserved G4 motif in the first exon of the MTOR gene as a model, we showed remarkable trinucleotide symmetry preservation across primates and broader mammals, with functional G4 regions displaying locally elevated symmetry relative to the codon-biased exonic background. Analysis of experimentally validated oncogenic G4s, including c-MYC, BCL2, VEGF, and KRAS, revealed that mirror and reverse complement symmetries converge around biologically important G4s. To quantify this feature, we formulated two complementary descriptors: the mirror symmetry index (MSI) and its non-palindromic variant (nMSI). Across 14 oncogene-promoter wild-type G4s, the majority scored MSI [≥] 0.80 (mean 0.884), with only the loop-rich ATG7, BCR, and MDM2 motifs falling below this value, and the KRAS promoter G4 reached individual significance against its mononucleotide-preserving null distribution (p = 0.042). Most decisively, each wild-type G4 scored higher on MSI than its experimentally confirmed G4-abolished mutant in 12 of 14 paired comparisons (sign test, p = 0.0065; mean {Delta}MSI = +0.089, mean {Delta}nMSI = +0.192); the two reversals (BCL2 and HIF-1) are attributable to scrambled mutant controls that introduce more balanced trinucleotide compositions rather than to failure of the index. The directional trend was reproduced across three independently published datasets, with nMSI [≥] 0.50 separating G4-forming from non-G4 sequences at 77.8% sensitivity and 100% specificity, although the collective per-sequence signal from mononucleotide-preserving shuffles remained a non-significant trend (Stouffer combined Z = 1.197, p = 0.116). This first report of trinucleotide symmetry in G4 motifs posits that coordinated nucleotide insertion and quadruplet maintenance act as an evolutionary forcing mechanism that pre-organizes single strands for G4 folding.

5
Building computational benchmarks: an Omnibenchmark reimplementation of a single-cell preprocessing pipeline evaluation

Choudhury, A.; Kitak, T.; Carrillo, B.; Busch, P.; Emons, M.; Gunz, S.; Koderman, M.; Luo, S.; Mallona, I.; Meara, A.; Wissel, D.; Robinson, M. D.

2026-05-05 bioinformatics 10.64898/2026.05.01.722166 medRxiv
Top 0.1%
12.7%
Show abstract

In the past few years, we have seen a veritable surge in single-cell (e.g., RNA sequencing) techniques and datasets, enabling increasingly detailed characterization of cellular heterogeneity across tissues and conditions. This surge in single-cell techniques has been complemented by a large number of analysis frameworks and pipelines, and a large parameter space and researcher degrees of freedom to use them. Many neutral benchmarks have been presented for various computational tasks, but most make design decisions that render them incompatible with each other, e.g., different datasets and metrics, or parameter sets used. In this work, we showcase a recently developed framework, Omnibenchmark, to build reproducible, extensible and standardized method comparisons. This not only facilitates the broad investigation of pipelines used in single-cell data analysis, but also highlights how the process of building benchmarks can be streamlined and unified. We do this as an initial proof-of-principle for an arms-length benchmark that evaluates five single-cell RNA sequencing pipelines (filtering to normalization to dimensionality reduction to clustering) on three datasets. This standardization enables benchmarks to be easily extended in several directions, including broader parameter sweeps, comparisons across software versions and architectures, isolation of pipeline steps, and integration of additional pipelines, datasets, and metrics.

6
BacNeMu: neutral mutation spectra reconstruction pipeline for bacteria

Skudnov, A.; Badamshin, E.; Efimenko, B.; Popadin, K.; Gunbin, K.; Denisov, S.

2026-07-02 bioinformatics 10.64898/2026.06.30.735404 medRxiv
Top 0.1%
12.6%
Show abstract

The mutational spectrum is an increasingly important molecular phenotype that quantitatively describes mutagenesis in a given gene and species, enabling future comparative analyses to reveal differences in underlying mutagenic processes, whether internal, such as DNA repair processes, or external, such as ecological niches and conditions. Mutation accumulation experiments, although time-consuming and costly, remain the standard approach for reconstructing bacterial neutral mutation spectra. Here, we present BacNeMu, a phylogenetically informed pipeline that reconstructs neutral mutational spectra of bacterial genomes using open databases GTDB, AnnoTree and KEGG Orthology, building on previously developed NeMu pipeline. BacNeMu reconstructs mutation spectra that closely match mutation accumulation experiments results while requiring substantially less time, enabling comparative analyses across diverse bacterial taxa. Applied to obligate aerobes and anaerobes, BacNeMu recovered the expected excess of T:A>C:G transitions, consistent with oxidative-damage-associated mutational patterns previously described in mitochondrial genomes and yeast single-strand. We further asked if any other ecologic factors influence a mutational spectrum. As a pilot we compared three species living under different temperatures: one strong thermophile - Thermotoga maritima, one psychrophile - Clostridium algidicarnis, and one with intermediate temperature tolerance - Psychrobacter sanguinis. In the thermophile, the relative frequency of T:A>C:G substitutions was higher than in the psychrophile, consistent with the hypothesis that GC-biased mutagenesis contributes to thermal adaptation, although C:G>T:A transitions predominate across all three species. BacNeMu provides a rapid, phylogenetically informed framework for generating biologically meaningful mutation spectra from open databases.

7
Systematic Evaluation of Feature Representations for Cancer-Associated sORF Prediction in Non-coding RNA

Rodrigues de Goes, F.; Mazheke, M.; Piveta Schnepper, A.; Karmakar, A.; de Souza, N.; Carvalho, R. F.; Basham, M.; Rossi Paschoal, A.

2026-06-20 bioinformatics 10.64898/2026.06.16.732659 medRxiv
Top 0.1%
12.5%
Show abstract

Short open reading frames (sORFs) within non-coding RNAs (ncRNAs) have arisen as a hidden layer of gene regulation, encoding small peptides that represent a new class of cancer regulators with diagnostic and therapeutic potential. However, inferring associations between sORFs to specific cancer types remains challenging and requires computational approaches for accurate prediction. Recently, the CoraL framework introduced the first computational approach for predicting cancer-associated peptides, focusing primarily on model architecture while overlooking how feature extraction strategies influence predictive accuracy. We present a systematic evaluation of machine learning models and feature extraction approaches to predict cancer-associated sORFs across 15 cancer types. We benchmarked seven traditional machine learning algorithms combined with three feature extraction methods: k-mer frequency, Word2Vec embeddings, and genomic language model (gLM)-based embeddings. To our knowledge, this is the first study applying gLM-derived embeddings to the prediction of cancer-associated sORFs in ncRNA. Our results show that traditional machine learning models with appropriate feature extraction outperform the CoraL baseline across all cancer types, achieving up to 10% higher accuracy in some of the 15 evaluated datasets. Interestingly, k-mer features consistently outperformed gLM embeddings without fine-tuning, suggesting that local sequence composition may provide more discriminative information for this task and that pre-trained genomic representations may require task-specific adaptation to fully capture these patterns. Additionally, we observed that the way sequences are tokenized, such as the k-mer length, can affect performance: longer fragments (e.g., k=7) sometimes reduced accuracy for Random Forest but had a smaller effect on MLP. Our findings suggest that appropriate feature engineering can provide greater improvements than increasing model complexity.

8
MKMC enables reference-free transcriptomic analysis using k-mer representations

Mboning, L.; Dlugosz, M.; Kokot, M.; Chen, J.; Costa, E. K.; Wu, M.-R.; Wang, S.; Bouchard, L.-S.; Deorowicz, S.; Pellegrini, M.

2026-07-10 bioinformatics 10.64898/2026.07.06.736868 medRxiv
Top 0.1%
12.5%
Show abstract

Traditional RNA-seq analysis depends heavily on genome alignment and gene annotation, limiting its utility in non-model organisms and introducing biases that can obscure regulatory complexity. We present MKMC (Multi-sample Kmer Counter), a scalable, reference-free toolkit for RNA-seq analysis that leverages k-mer-based statistics to detect biological variation without requiring alignment. MKMC integrates fast k-mer counting, abundance matrix generation, normalization, dimensionality reduction, and differential analysis into a unified workflow. Across diverse datasets, MKMC recapitulates key biological signals--including sex differences in killifish liver--and matches alignment-based pipelines in differential expression analysis and transcriptomic age prediction. Notably, MKMC detects isoform-specific events missed by traditional methods, one of which we validated using in situ hybridization. These results reveal previously hidden isoform-level regulatory events that contribute to sex-and age-associated transcriptional programs. MKMC offers a robust, extensible alternative to alignment-based approaches, enabling transcriptomic discovery across both model and non-model systems. While we focus here on RNA-seq as a primary application, MKMC is broadly applicable to any k-mer-based analysis of next-generation sequencing data.

9
AbSolution: interactive exploration of sequence-derived features in AIRR-seq repertoires

Garcia-Valiente, R.; Triantafyllou, C.; van Schaik, B.; Jongejan, A.; Pollastro, S.; Anang, D. C.; Guikema, J. E.; de Vries, N.; Hoefsloot, H. C.; van Kampen, A. H. C.

2026-05-22 bioinformatics 10.64898/2026.05.20.726477 medRxiv
Top 0.1%
12.4%
Show abstract

High-throughput sequencing of B-cell and T-cell immune receptor repertoires provides unprecedented insight into adaptive immune responses. The data produced are structured by clonal relationships and somatic mutation signatures, and yield extremely rich information in sequence-derived features, including physicochemical properties and compositional patterns. However, integrated analysis across datasets, conditions, and time points remains challenging. Current analytical tools typically focus only on certain features within individual repertoires, without enabling integrated, multivariable comparisons across datasets, conditions, and time points to address their diversity and variability. Here we present AbSolution, a user-friendly and flexible interactive application for comprehensive exploration of immune repertoires and their sequence-based properties. AbSolution enables multiscale analysis of thousands of sequence-derived features across receptor regions, while accounting for V(D)J usage, clonal composition and experimental groupings. We demonstrate its utility by identifying distinct sequence-based profiles associated with dominant (highly abundant) and non-dominant B-cell clones in peripheral blood BCR repertoires from patients with idiopathic inflammatory myopathies, and with antigen-responsive T-cell populations over time in a longitudinal in vitro antigen-stimulation dataset. Through interactive, interlinked visualizations, statistical feature selection and multi-sample comparisons, AbSolution facilitates integrated feature profiling that supports the interpretation of immune selection processes and enables systematic analysis of complex repertoire datasets.

10
Hidden sampling biases inflate performance in gene regulatory network inference

Stock, M.; Ratajczak, F.; Bertin, P.; Hoermanseder, E.; Bengio, Y.; Hartford, J.; Falter-Braun, P.; Heinig, M.; Tong, A.; Scialdone, A.

2026-07-14 bioinformatics 10.64898/2025.12.19.695616 medRxiv
Top 0.1%
12.3%
Show abstract

Accurate reconstruction of gene regulatory networks (GRNs) from single-cell transcriptomic data remains a major methodological challenge. Recent machine learning approaches, particularly graph neural networks and graph autoencoders, have reported improved performance, yet these gains do not consistently translate to realistic biological settings. Here, we show that a key reason for that is the way negative regulatory interactions are sampled for supervised training and evaluation. We find that widely used sampling strategies introduce node-degree biases that allow models to exploit trivial graph-structural cues rather than biological signals. Across multiple benchmarks, simple degree-based heuristics match or exceed state-of-the-art graph neural network models under these biased evaluation protocols. We further introduce a degree-aware sampling approach that eliminates these artifacts and provides more reliable assessments of GRN inference methods. Our results call for standardized, bias-aware benchmarking practices to ensure meaningful progress in supervised GRN inference from single-cell RNA-seq data.

11
DDTRN: Predicting Bacterial Transcriptional Regulatory Networks Based on Gene Sequences using Dual Descriptor

Nie, P.; Ma, B.-G.

2026-07-01 bioinformatics 10.64898/2026.06.30.735580 medRxiv
Top 0.1%
12.3%
Show abstract

Accurate computational reconstruction of bacterial transcriptional regulatory network (TRN) from sequence information alone remains a fundamental challenge in systems biology, particularly for non-model organisms lacking extensive transcriptomic data. We present DDTRN, a sequence-driven framework that formulates TRN inference as a binary classification task over concatenated regulator-target gene sequence pairs and employs a Dual Descriptor (DD) model to predict regulatory interactions. The DD architecture represents a sequence into two learnable components: Composition Weight Map (CWM) and Position Weight Function (PWF). We comprehensively evaluate DDTRN against six conventional machine learning baselines across eight benchmark bacterial datasets, including E. coli (DREAM5, RegulonDB), B. subtilis, S. enterica, C. glutamicum, M. tuberculosis, P. aeruginosa, and S. coelicolor. DDTRN achieves superior overall performance, attaining average AUROC and AUPR scores of 0.869 and 0.868, respectively, with particularly pronounced advantages at lower descriptor ranks where positional weighting compensates for limited sequence context. Systematic sensitivity analyses of rank, embedding dimension, and basis function count reveal stable optimal operating regimes, while subsampling experiments demonstrate strong robustness even with limited training data. Interpretability analyses show that PWF learns distinct periodic contributions across different rank granularities and that CWM preferentially weights meaningful k-mers. A case study on E. coli dataset further illustrates that DDTRN identifies method-specific candidate targets complementary to those proposed by conventional approaches. By operating solely on genomic sequence, DDTRN provides a scalable, interpretable, and data-efficient framework for bacterial TRN inference in species where expression data are scarce, and it establishes a foundation for future multimodal integration with condition-specific regulatory information.

12
Entropy Fusion DNA: Alignment-Free Gene Fusion Detection through Entropy and Mutual Information Descriptors

Benevento, G.; Malandrino, D.; Ture, A.; Zaccagnino, R.

2026-05-30 bioinformatics 10.64898/2026.05.27.728176 medRxiv
Top 0.1%
12.0%
Show abstract

Gene fusions are clinically relevant genomic alterations and key cancer biomarkers. Their computational detection remains dominated by alignment-based pipelines, whose reliance on read mapping, reference annotations, and heuristic filtering makes them sensitive to mapping ambiguities, annotation incompleteness, repetitive regions, and false positives. Recent machine learning (ML) strategies aim to learn fusion-related patterns directly from sequencing data, but their adoption is still limited by dataset-specific biases, synthetic data artifacts, class imbalance, and representations that may overlook the structural organization of biological sequences. Theoretical and statistical sequence descriptors remain underexplored as efficient tools for capturing informative structural signals in biological reads. In this work, we investigate whether fusion-related information can be inferred directly from the statistical organization of DNA sequences. Each sequence is encoded into a compact, interpretable, and alignment-free feature space combining Shannon and Renyi entropy, lagged and base-resolved mutual information, GC content, and rarefied k-mer richness descriptors. Our goal is to assess whether these information-theoretic features encode discriminative sequence signatures associated with fusion events. For discriminating fusion-derived from non-fusion sequences, nested cross-validation selected K-nearest neighbors as the most effective classifier, achieving strong held-out performance on the balanced benchmark (AUROC = 0.892, AUPRC = 0.865). The same representation was then evaluated on fusion-positive samples for fusion partner prediction and breakpoint localization, achieving strong top-k partner identification accuracy and stable breakpoint regression performance. Moreover, a two-stage strategy in which the binary classifier first filters candidate reads further improved partner prediction, suggesting its use as an enrichment step for downstream fusion characterization. Although performance decreased under repeated fusion-pair-disjoint evaluation, it remained clearly above random expectation, supporting the transferability of the proposed descriptors to unseen fusion pairs. Breakpoint-centered validation further revealed increased local sequence complexity, altered short-range dependency structure, and modest but significant microhomology enrichment around fusion regions. Such findings support an interpretable alignment-free framework where information-theoretic features provide predictive and biologically informative signals for gene fusion analysis. The framework is available at: https://github.com/FLaTNNBio/EntropyFusionDNA Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=73 SRC="FIGDIR/small/728176v1_ufig1.gif" ALT="Figure 1"> View larger version (23K): org.highwire.dtl.DTLVardef@805fa3org.highwire.dtl.DTLVardef@6f6cdorg.highwire.dtl.DTLVardef@1352c8borg.highwire.dtl.DTLVardef@1ff780b_HPS_FORMAT_FIGEXP M_FIG C_FIG HighlightsO_LIAlignment-free information-theoretic DNA descriptors detect gene fusions. C_LIO_LIResolved mutual-information features provide the strongest predictive signal. C_LIO_LITwo-stage screening enriches partner-gene prediction and breakpoint analysis. C_LI

13
Scanning transcriptomes for nonlinear, domain-level similarities using hmSEEKR

Li, S.; Sprague, D. A.; Eberhard, Q. E.; Boyson, S. P.; Laederach, A.; Calabrese, J. M.

2026-07-08 bioinformatics 10.64898/2026.07.03.736302 medRxiv
Top 0.1%
11.9%
Show abstract

Long noncoding RNAs (lncRNAs) play roles in gene regulation across kingdoms of life. However, lncRNAs with related functions often lack linear sequence similarity, making it difficult to leverage studies of one lncRNA to inform the understanding of others. We describe a k-mer-based hidden Markov model, hmSEEKR, that enables the scanning of transcriptomes for regions of non-linear sequence similarity to a query domain, without prior knowledge of where within the transcriptome the similarities may be located. When individual lncRNA domains were used as search features, hmSEEKR successfully identified regions in other RNAs that harbor non-linear sequence similarity and bind similar sets of proteins. Applying hmSEEKR to transcriptome-wide searches, we found that certain domains within the lncRNAs XIST, NEAT1, and MALAT1 exhibited widespread regional similarity to both lncRNA and protein-coding genes, while others were more unique, exhibiting similarity to ~100 genes or fewer. Combinatorial searches uncovered RNAs containing sequential matches to core functional domains of XIST and NEAT1, and eCLIP-inferred protein-interaction networks within these RNAs more closely resembled those of XIST and NEAT1, respectively, than would be expected by chance, suggesting the searches recovered RNAs with similar biological properties. Finally, within annotated sets of cis-activating and cis-repressive lncRNAs, we observed opposing enrichments for similarity to domains associated with transcription-promoting complexes and heterogeneous nuclear ribonucleoprotein (hnRNP) binding, respectively, suggesting the enriched sequences may contribute to regulatory functions. hmSEEKR can be applied with minimal training data and enables the a priori discovery of RNA domains that share nonlinear similarity, offering a sequence-informed approach to discover functional elements within noncoding transcriptomes.

14
HeartBioPortal 3.0: an integrated cardiovascular genomics knowledge environment for molecular, clinical and population-scale interpretation

Vand, K.; Badia, N.; Khomtchouk, B.; Janga, S. C.

2026-07-01 cardiovascular medicine 10.64898/2026.06.28.26356792 medRxiv
Top 0.1%
11.8%
Show abstract

Cardiovascular genomics is producing rapidly expanding genetic, molecular, phenotypic and clinical data, yet relevant evidence remains fragmented across resources and difficult to translate into actionable biological and ultimately translational knowledge. HeartBioPortal (HBP) is a browser-based cardiovascular knowledge environment that was developed to address this problem by organizing omics, variant, phenotype and clinical evidence centered around gene queries. Here we describe HBP 3.0, a major update that expands both the data architecture and interpretive interface. This update introduces DataHub, a reproducible data-engineering layer for source ingestion, standardization, variant-centered aggregation, provenance tracking and compact serving artifacts. The release integrates cardiovascular clinical practice guideline context through a graph-backed clinical knowledge layer; incorporates cardiovascular summary statistics from the Million Veteran Program and public aggregate resources; expands source-preserving population frequency, variant annotation and structural-variant; and adds gene profile, drug-discovery and protein-context layers. HBP 3.0 incorporates 594.3 million allele-frequency observations across 18.1 million rsIDs, 3.04 million exon-enriched structural-variant records, 66.9 thousand protein isoforms with 3.26 million non-exon protein feature annotations, 17,128 gene-drug records, and a clinical guideline knowledge graph with 42,895 entities and 106,304 relationships. The redesigned gene dossier view combines phenotype filtering, annotation composition, persistent selected-detail panels and exportable chart data in one workflow. HBP 3.0 is designed to help cardiovascular and eventually cardiometabolic researchers move from a genetic or genomic signal to biological knowledge and potentially clinical and therapeutic context while preserving source provenance and interpretive boundaries. Database URL: https://www.heartbioportal.com/

15
AMaNITA: an end-to-end workflow for native tRNA nanopore sequencing data analysis

Katopodi, X.-L.; Pryszcz, L. P.; Llovera, L.; Ollivier, A.; Cozzuto, L.; Ponomarenko, J.; Novoa, E. M.

2026-06-17 bioinformatics 10.64898/2026.06.16.732588 medRxiv
Top 0.1%
11.8%
Show abstract

Transfer RNA (tRNA) molecules serve as essential adapters during protein translation. While direct RNA sequencing (DRS) via Oxford Nanopore Technologies has emerged as a powerful platform for systematic tRNAome profiling, we currently lack a simple and robust statistical framework for nanopore tRNA data analyses. Here, we address this gap by developing AMaNITA (Abundance, Modifications, and Nanopore Intensity Toolbox Application), an end-to-end bioinformatic workflow that enables simplified, robust, and scalable analyses of nanopore native tRNA sequencing datasets. AMaNITA streamlines the entire analytical trajectory: from upstream processing (basecalling, mapping, filtering, batch effect correction) to downstream assessment of differential tRNA abundance and modification stoichiometry. The workflow generates an interactive HTML report for data exploration and analysis, allowing the user to download the source data files and resulting plots. AMaNITA can be executed using Singularity from the command line, without requiring installation of dependencies.

16
Somatic variant detection in normal tissues from single-cell sequencing data

Luo, R.; Wang, Z.; Dou, J.; Bhamidipati, S. V.; Kalra, D.; Grochowski, C. M.; Doddapaneni, H. V.; Gibbs, R. A.; Chen, K.; Chen, R.

2026-06-14 bioinformatics 10.64898/2026.06.10.731451 medRxiv
Top 0.1%
11.8%
Show abstract

A crucial advantage of single-cell sequencing (SCS) is its ability to identify somatic variants in individual cells, enabling phylogenetic analysis of cellular populations within bulk tissues. While identifying somatic variants in tumor tissues via SCS has become a common practice, doing so in normal tissues remains challenging due to the rarity of somatic variants in normal cells. To evaluate the feasibility of somatic variant calling from widely available single-nucleus RNA-seq (snRNA-seq) and single-nucleus ATAC-seq (snATAC-seq) data, we profiled a Cell-line mix of six HapMap samples prepared by the SMaHT consortium using 10x Genomics 5 snRNA-seq (12k cells with 36k mean reads per cell) and snATAC-seq (11k cells with 14k median high-quality fragments per cell) for variant calling. PacBio long-read whole genome sequencing (WGS) data (109x) generated from individual cell lines were used as ground truth. Two computational tools, Monopogen and SComatic, were used for somatic variant calling from the SCS data. Monopogen achieved single nucleotide variant (SNV) detection accuracies of 93.30% in the snRNA-seq and 99.64% in the snATAC-seq data, both of which outperformed SComatic (74.35% and 94.29%, respectively). Monopogen also consistently detected somatic SNVs at cellular fractions as low as 0.5% (2.54% in snRNA and 0.81% in snATAC) in individual samples. Notably, snATAC-seq exhibited higher genomic coverage breadth and larger number of variants detected than snRNA-seq. While the SCS data have lower overall genome coverage than that of the bulk WGS, the single-cell level variant resolution allows Monopogen to assign variants to their cells of origin with over 80% accuracy in both RNA and ATAC modalities, thereby facilitating studies of clonal evolution and cell-type-specific mutagenesis. Other benchmarking methods were also evaluated (DeepVariant, Cellsnp-lite and Mutect2) for comparison. In conclusion, our study demonstrated the feasibility of performing reliable single-cell somatic mutation calling in a cell-line mixture and discussed the strengths and limitations of current computational methods when applied to normal tissues.

17
DanioDecima: A DNA sequence-to-function model of zebrafish embryogenesis

Voges, M. J.; Kim, Y. J.; Frank, M.; Iovino, B.; Senbabaoglu, Y.; Royer, L. A.

2026-05-31 genomics 10.64898/2026.05.29.728876 medRxiv
Top 0.2%
11.7%
Show abstract

Deep learning DNA sequence-to-function models offer the promise of gaining mechanistic insights into genome regulation, however their performance is often limited by data scarcity in the species of interest. We present DanioDecima, a zebrafish-specific model leveraging transfer learning from human and mouse-trained models to predict tissue- and cell-type-specific gene expression during zebrafish embryogenesis. Initializing DanioDecima with pretrained human and mouse Borzoi and Decima weights raises the median pseudobulk Pearson r sub-stantially across cell-types and improves gene-level correlations of test set genes. An in silico directed-evolution loop guided by DanioDecima scoring generated synthetic promoters whose motif architectures cluster by the expected target lineage. These findings exemplify a cross-species transfer learning methodology for sequence-to-function models, and position DanioDecima as a practical resource for zebrafish regulatory engineering.

18
TransXplorer: An automated translational discovery platform for RNA-seq data

Verma, V. M.; Oler, E.; Syed, H.; Han, S.; Berjanskii, M.; Mason, A. L.; Wishart, D. S.; Wong, G. K.-S.

2026-05-16 bioinformatics 10.64898/2026.05.15.724657 medRxiv
Top 0.2%
11.6%
Show abstract

RNA-seq experiments routinely identify thousands of differentially expressed genes, but translating these into biological insights and therapeutic hypotheses often requires integrating multiple tools. Existing web platforms such as iDEP, NetworkAnalyst, and GEPIA2 address individual steps, differential expression, network visualization, or TCGA queries, but lack a unified environment spanning raw data processing to clinical and pharmacological interpretation. TransXplorer (https://transxplorer.org) is a freely available web platform that addresses this limitation by integrating the complete RNA-seq analytical workflow. It supports processing from raw FASTQ files using HISAT2 or Salmon, as well as direct GEO dataset import with automated metadata handling. Differential expression analysis is implemented via DESeq2, edgeR, and limma-voom, followed by functional enrichment across more than 1,800 species using Bioconductor resources. Batch effects are automatically detected and corrected using a composite of PVCA, kBET, and Silhouette metrics without requiring predefined batch annotations. Downstream analyses include co-expression network construction (WGCNA), protein-protein interaction mapping (STRING), cell-type deconvolution, and transcription factor inference using integrated DoRothEA and TFLink resources. The platform further links gene signatures to drug candidates through DGIdb and OpenTargets and enables survival and tumour-normal comparisons across TCGA cohorts. Application to cardiac endothelial differentiation (GSE151427) and kidney renal papillary cell carcinoma (TCGA-KIRP) datasets demonstrates accurate batch correction, biologically consistent pathway enrichment, recovery of expected cell-type proportions, and identification of clinically relevant genes and drug candidates. TransXplorer is freely available without a login.

19
Estimation of splicing metrics for NMD-sensitive transcripts

Zavileyskiy, L.; Vlasenok, M.; Kuznetsova, A.; Skvortsov, D. A.; Pervouchine, D. D.

2026-07-04 bioinformatics 10.64898/2026.06.30.735642 medRxiv
Top 0.2%
11.6%
Show abstract

Alternative splicing is commonly quantified using the Percent-Spliced-In (PSI) metric, which measures the relative abundances of alternatively spliced isoforms. However, some transcript isoforms are targeted by the nonsense-mediated decay (NMD) pathway, introducing a strong bias that leads to underestimation of their true splicing rates. To correct for this bias, we developed an analytical framework and a set of statistical models employing a linear fractional transformation depending on a single parameter capturing the degradation rate of NMD-sensitive transcripts relative to normal mRNA decay. Using Gaussian mixture models, we demonstrated a clear separation of splicing events into two classes, responders and non-responders, with the former exhibiting strong upregulation upon NMD inhibition and the latter showing little or no response. Moreover, non-responders displayed higher coding potential and stronger translation signals both upstream and downstream of the stop codon, which are characteristic of NMD escape through translational readthrough. We further showed that incorporation of event-specific relative decay rates improves the interpretation of differential splicing patterns for NMD-sensitive transcripts. In sum, our results provide a solid framework for unbiased estimation of splicing metrics in NMD-sensitive transcripts from short-read RNA-seq data, without requiring NMD inhibition experiments.

20
Robust data-driven gene expression inference for RNA-seq using curated intergenic regions

Brandulas Cammarata, A.; Fonseca Costa, S. S.; Rosikiewicz, M.; Roux, J.; Wollbrett, J.; Bastian, F. B.; Robinson-Rechavi, M.

2026-05-20 genomics 10.1101/2022.03.31.486555 medRxiv
Top 0.2%
11.3%
Show abstract

RNA-Seq is a powerful technique to provide quantitative information on gene expression. While many applications focus on measuring expression levels, accurately distinguishing between actively and inactively transcribed genes is equally important for understanding gene function, development, and disease mechanisms. However, setting a biologically meaningful threshold for calling genes expressed is challenging due to variability in noise levels across different protocols, experiments or biological samples. We propose to define this threshold per sample relative to the background level observed in inactive genomic features, inferred by the amount of reads mapped to intergenic regions of the genome, and to call genes expressed if their level of expression is significantly higher than the estimated background noise. This approach can be applied to a single RNA-Seq library as well as to a combination of libraries from the same condition, in model and non-model organisms. We show that our method yields a more accurate prediction of expression state than existing methods, illustrated by consistent expression calls for biological replicates in the same tissue.